Google2026-10-08 02:10:21Google open-sources EmbeddingGemma 2, a 567MB multimodal retrieval model built to run offline on phonesGoogle on Oct. 6 released EmbeddingGemma 2, its first native multimodal open-source embedding model, bringing text, code, image, video and audio retrieval into one system. The model has 740 million parameters and is designed to run locally across phones, laptops and even browsers, with the fully quantized multimodal version using about 567MB of memory on a Pixel 11 Pro, according to Google. The release expands Google’s earlier text-only embedding work by mapping multiple content types into a shared 768-dimensional space. That allows a text query such as “birds singing” to retrieve both bird photos and bird-call recordings, and it also enables voice-to-video matching and mixed-content retrieval across product pages, media libraries and local files. Google said the model supports an 8K context window, more than 100 languages, and larger single-input capacities including about 5.5 minutes of audio, 29 images or 58 video frames. Google’s published benchmarks showed EmbeddingGemma 2 ahead of Jina v5 Omni-Nano in image and video retrieval, while code retrieval on MTEB Code improved from 68.76 to 78.68 versus the previous generation. The company also pointed to local-first use cases, including offline RAG with Gemma 4, on-device media search, and codebase indexing without sending files to the cloud.20